Papers with RLVR-based attacks
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have gained significant attention due to their objective and verifiably verifier reward signals. |
| Approach: | They propose to exploit RLVR for alignment reversibility by using GRPO to reverse alignment with merely 64 harmful prompts without responses. |
| Outcome: | The proposed method outperforms fine-tuning and RLHF in reasoning and code generation tasks while maintaining general capabilities. |